Tag
15 articles
Security researchers have developed a new technique called 'context bombing' that prevents malicious AI agents from executing harmful actions by overwhelming them with excessive contextual data.
OpenAI's internal AI red-teaming model, GPT-Red, outperformed human red-teamers 84% to 13% on prompt injection tests and discovered novel attack techniques.
OpenAI introduces GPT-Red, an automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness. This innovation represents a significant advancement in AI safety research.
Learn how to implement defensive prompt injection techniques using context bombing to protect AI systems from malicious manipulation.
OpenAI introduces ChatGPT's Lockdown Mode to protect sensitive data from prompt injection attacks by disabling web access and research features.
Learn about prompt injection attacks and how OpenAI's new Lockdown Mode aims to protect sensitive data in AI systems.
A simple GitHub issue could have compromised Anthropic’s Claude Code action, exposing projects that use it to potential data breaches and unauthorized access.
A developer has revealed how a malicious code addition in the popular Java library jqwik could have instructed AI coding agents to delete application output, highlighting serious security vulnerabilities in AI-assisted development.
Google warns that malicious web pages are poisoning enterprise AI agents through indirect prompt injections, exploiting hidden HTML code to manipulate AI systems.
Learn how cybercriminals trick AI systems into leaking data and executing malicious code through subtle prompt injection attacks. Understand the risks and protection methods.
Security researcher Aonan Guan exploited prompt injection flaws in AI agents from Anthropic, Google, and Microsoft, stealing API keys. All three companies paid bug bounties but did not issue public advisories.
OpenAI reveals new defenses against prompt injection attacks and social engineering in ChatGPT, strengthening AI agent security through constrained workflows and enhanced data protection.